Annals of Internal Medicine
● American College of Physicians
Preprints posted in the last 30 days, ranked by how well they match Annals of Internal Medicine's content profile, based on 28 papers previously published here. The average preprint has a 0.02% match score for this journal, so anything above that is already an above-average fit.
Adapa, K.; Mosaly, P. R.; Yu, F.; Moore, C.; McGurk, R.; Das, S.; Mazur, L.
Show abstract
Radiation oncology has a long history of developing in-house health information technology (HIT) tools such as quality assurance (QA) checklists, yet there is little guidance from professional bodies on how to implement these tools in complex clinical environments. Building on our previous work that used human-centered participatory co-design, the Task-User-Representation-Function (TURF) framework, and multi-method usability evaluations to design and develop an enhanced dosimetry QA checklist (DQC), this study investigated the barriers and facilitators (determinants) to implementing the enhanced DQC in a radiation oncology clinic, examined implementation strategies, proposed an implementation framework for QA checklists in radiation oncology, and assessed four implementation outcomes: acceptability, appropriateness, feasibility, and adoption. We conducted a qualitative implementation study using an abductive research approach at an academic medical center. All key stakeholders (dosimetrists, physicists, trainees, and software developers) participated in semi-structured interviews, field observations, and surveys across pre-implementation, implementation, and post-implementation phases. Data were analyzed using a hybrid inductive-deductive approach, with deductive coding guided by an adapted Consolidated Framework for Implementation Research (CFIR) mapped to the Unified Theory of Acceptance and Use of Technology and by the Expert Recommendations for Implementing Change (ERIC) compilation. We identified 4 CFIR constructs and 12 sub-constructs as barriers, with structural characteristics and planning showing the highest negative valence, and 5 CFIR constructs and 19 sub-constructs as facilitators, with relative advantage, culture, and leadership engagement showing the highest positive valence. Participants' suggestions mapped to 19 ERIC strategies in 7 clusters, and the CFIR-ERIC matching tool identified 14 evidence-based strategies in 4 clusters that informed a proposed phased implementation framework. Acceptability, appropriateness, and feasibility scores improved significantly from pre-implementation to implementation for all professional roles (p<0.05), yet adoption reached 100% only in the sixth week of implementation. These findings highlight the value of combining subjective and objective implementation outcomes and provide a practical, evidence-based framework for implementing in-house QA checklists in radiation oncology that warrants validation in diverse settings.
Joseph, A.; Kearney, K.; Henricks, C.; Morgan, J. L.; Tan, W.; Shafer, K.; Wrobel, C.; Lacelle, C.; Burns, K.; Jawaid, A.; Tapaskar, N.; Solmonson, A.; Nelson, D. B.; Truby, L. K.
Show abstract
Background: Adult congenital heart disease (ACHD) patients are prone to HLA-antibody formation from multiple surgeries, transfusions, and prosthetic surgical material. Females with ACHD may accrue additional, non-surgical alloantigen exposure. Whether sex modifies the impact of allosensitization on heart transplant (HT) access and outcomes in ACHD remains unknown. Methods: We retrospectively analyzed the OPTN/UNOS registry of adults with ACHD listed for first-time HT (2018-2025). Sensitization was defined by calculated panel reactive antibodies (cPRA) at listing. We tested the sex x sensitization (highly sensitized, cPRA >50%) interaction on transplant access using Fine-Gray competing-risks regression, treating transplantation as the event of interest and death or removal from the waitlist as competing events, and on post-transplant survival using multivariable Cox proportional-hazards regression, both adjusted for age at listing, mechanical support at listing, and the number of distinct prior cardiac surgery categories. Results: Among 856 candidates (38% female), females were more often highly sensitized than males (23% vs 14%; age-adjusted OR 1.81, 95% CI 1.26-2.61), even after adjusting for surgical burden. Sensitization reduced transplant access in females (84% to 71%; median wait 60 to 110 days, p < 0.001) but not males (79% vs 79%, median wait 88 vs 98 days). In adjusted Fine-Gray models, the subdistribution hazard for transplant was reduced in sensitized females (sHR 0.54, 95% CI 0.41-0.72) with no effect in males (sHR 0.96, 95% CI 0.73-1.26), and the sex x sensitization interaction was significant (interaction sHR 0.64, 95% CI 0.44-0.94, p = 0.02). Post-transplant mortality was numerically higher in sensitized than non-sensitized candidates in both sexes and the sex x sensitization interaction on 1-year mortality was not significant. The sex-asymmetric effect persisted and was more pronounced in the multiorgan candidates. Conclusions: Allosensitization is not a sex-neutral barrier to transplant in HT candidates with ACHD. Females are more sensitized and have reduced transplant access without differences in 1-year mortality. The female excess in sensitization is not accounted for by surgical burden, and the exposures responsible remain to be defined. These findings warrant a sex-aware listing strategy and further studies.
McLean, K. W.; LaBonte, J.; Macaulay, K.; Kassam-Adams, S.
Show abstract
This study documents the derivation and validation of a deterministic algorithm for cause-of-death (COD) ascertainment from longitudinal real-world medical claims data, evaluated against an independent state-level death certificate file. Death certificates are the dominant reference standard in mortality research but carry well-documented limitations, including primary-cause error rates estimated at 20-40\% across empirical studies. A matched analytic cohort of 216,382 individuals (Connecticut death records, 2017--2025, age 25 and above) was constructed after exclusion of mechanism-of-injury cases and removal of ill-defined symptom-code entries from both sources. Concordance between algorithmic and certificate-based COD was assessed through three complementary frameworks: age-stratified positive predictive value (PPV) at the ICD-10-CM chapter level under a full-set concordance scenario; mean absolute rank difference (MARD) for chapters identified by both sources; and analyses of breadth, depth, and code-level specificity of COD reporting. Chapter-level PPV was strongest for individuals aged 55 and above, with all estimates representing conservative lower bounds given the known error rate of the certificate reference standard. The algorithm consistently reported broader and more granular contributing cause profiles than the death certificate, with discordances directionally consistent with the well-documented tendency of certificates to under-report contributing conditions. These findings support the conclusion that algorithmic COD ascertainment from longitudinal claims data is a feasible and scalable alternative to certificate-based attribution and, at population scale, a principled methodology for characterising death certificate error rates beyond what small-sample chart review studies can achieve.
Qian, Z.; Khera, A.; Makhnoon, S.; Chapman, B. E.; Bryant, B.; Sayers, M.; Compton, F.; Eason, S.; Xing, C.; Ahmad, Z.
Show abstract
Background. Cardiovascular-kidney-metabolic (CKM) syndrome affects nearly 90% of US adults, yet most individuals at early, modifiable stages remain unidentified outside clinical care. Blood donation centers offer a scalable, non-clinical venue for CKM screening, but the potential benefit of screening in this context remains unclear. We projected the population-level impact of effective digital return of results (ROR) to inform the design of a pragmatic trial. Methods. We developed a Monte Carlo simulation (100,000 iterations) of the incident major adverse cardiovascular events (MACE), end-stage renal disease (ESRD), and type 2 diabetes (T2DM) preventable by ROR-prompted, guideline-concordant follow-up among donors in CKM Stages 1-2. The estimand counts only events averted by donors who act because of ROR; the intervention effect was modeled directly on strictly positive support, and action was translated into prevented events through a hazard-based cumulative-incidence difference that counts each donor at most once. We evaluated 18 design cells (donor volumes 300,000, 1 million, and 8 million/year; 5- and 10-year horizons; action-rate gains of +10, +20, and +30 percentage points [pp]) and, in a complementary two-arm simulation, the assurance (expected power) of detecting the effect in a single deployment. Results. Under the primary +20 pp scenario, ROR at a single large blood center (300,000 donors/year) is projected to prevent a median of 2,201 events (95% uncertainty interval [UI], 1,099-4,364) over 10 years, scaling to 58,526 (29,154-116,769) at the national donor pool. All 18 design cells had strictly positive 95% lower bounds. The number needed to screen was 136 and the screening cost $2,045 per event prevented (at $15/donor), both invariant to donor volume. Impact scaled linearly with volume and effect size but sub-linearly with the horizon. Detection of the effect was effectively certain at gains of +20 pp or larger (assurance [≥]99.6% in every cell and >99.9% in all but the smallest 5-year cell). Conclusions. Even under the conservative scenario, digital CKM ROR at blood donation centers is projected to prevent hundreds to tens of thousands of incident cardiometabolic events at a screening cost per event well within accepted prevention benchmarks, providing prospective, quantitative justification for a pragmatic, randomized evaluation of digital ROR in non-clinical screening settings.
Mata-Robles, S.; Khalaf, K.; Kelley, J.; Chauhan, A.; Balian, L.; Linnes, J. C.; Rodriguez, N. M.
Show abstract
Point-of-care Hepatitis C Virus (HCV) RNA assays reduce diagnostic turnaround time but depend on benchtop instrumentation and continuous electricity, limiting their deployment in the harm reduction and community settings where confirmatory testing is most needed, as people who use drugs (PWUD) carry a disproportionate share of the HCV burden in the United States. This is a systemic problem in diagnostic development, where decision-making and design requirements overlook point-of-use stakeholders. Closing the gap requires integrating real-world constraints throughout design rather than validating against user needs once a product already exists. Here, we apply a human-centered design (HCD) approach to inform rigorous, stakeholder-derived design requirements, implementation considerations, and early value proposition for a novel point-of-need HCV RNA test intended for deployment in harm reduction and community settings in Indiana. To determine design specifications grounded in real-world context, our objectives were (1) identifying and characterizing context-specific experiences and barriers to HCV testing among higher-risk populations; (2) assessing the perceived benefits and acceptability of the proposed test within real-world settings across direct and indirect user groups; and (3) translating the user needs and contextual constraints into design requirements and implementation considerations that support the test's clinical, operational, and user-centered value. We conducted 18 semi-structured interviews with frontline staff and HCV testing/treatment pipeline experts (n=11) and people who get tested (n=7) across harm reduction organizations, syringe service programs, and community testing settings, analyzed using Rapid Qualitative Analysis guided by the PARRQA framework. Stakeholders responded positively to a single-encounter point-of-need RNA test, and implementation considerations, including funding restrictions, staffing structures, and diverse deployment settings, directly shaped design requirements spanning turnaround time, sample type and volume, portability, result output, target operator, and ease of use. Benchmarking these stakeholder-derived specifications against the FIND Dx HCV target product profile (TPP) showed that stakeholder input confirmed, modified, or extended several TPP criteria and introduced requirements the TPP does not address. Together, these objectives constitute an upstream, evidence-driven design process that translates contextual and stakeholder knowledge into actionable engineering requirements, highlighting the need for diverse stakeholder engagement at all stages of the design process for closing the translation gap between laboratory-validated diagnostic tools and effective point-of-need deployment.
Scherer, L. D.; Matlock, D. D.; Cronin, J.; Gritz, M.
Show abstract
Multi-Cancer Detection (MCD) tests can detect more than 50 different types of cancer using a blood test. Recently passed law in the U.S. guarantees that Medicare will pay for these tests when they are FDA approved and show evidence for clinical benefit. This manuscript provides estimates of the cost of MCD tests to Medicare under different assumptions of cost per test, eligibility, and screening uptake in the eligible population. This manuscript additionally estimates the cost of follow-up testing resulting from false positive results, which are considered avoidable costs caused by the screening test.
Bunning, B. J.; Weng, Y.; Wu, D. J.; Hui, G.; Hope, J. E.; Pandurangan, V.; Lopez, I.; Everett, S.; Chen, J. H.; Desai, M.
Show abstract
Doctors increasingly rely on AI in the clinic, yet which report features make AI-generated responses useful and trustworthy remains unclear. In this randomized mixed-methods study, 34 oncology physicians provided 294 ratings of four blinded AI systems across five vignettes, alongside 20 semi-structured interviews analyzed with a prespecified LLM-assisted qualitative pipeline. Despite similar references, an evidence-graded report adapted from OpenEvidence was rated significantly lower in overall utility than standard OpenEvidence (mean difference, -0.96; 95% CI, -1.26 to -0.66; P<.001). Qualitative analysis identified six themes and seven design requirements. Oncologists valued rapid orientation, evidence retrieval, and verification, preferring concise, scannable reports with quantitative outcomes, recognizable bolded guidelines, explicit uncertainty, and verifiable citations. Trust deteriorated with citation mismatch, buried provenance, evidence misclassification, overconfident recommendations, and poor organization. Evidence presented differently can alter perceptions of clinical utility and trust; accuracy alone is insufficient, and report design must also be empirically evaluated.
Bergman, H. I.; Liu, V. N.; Austin, B.; Sanghera, R.
Show abstract
Objectives Safety claims for ambient artificial intelligence (AI) scribes rest on automated judges that detect documentation errors and grade clinical risk. Expert reviewers are under-sensitive and disagree with one another, so no gold standard exists and validation cannot mean accuracy. We tested whether such judges are a defensible instrument: reproducible, within the envelope of expert disagreement, and non-differential across arms. Methods Pre-registered, blinded validation study nested in a multi-country simulation of ambient AI documentation (English setting), reported per GRRAS. Ten external clinicians independently adjudicated a stratified sample of 434 pipeline flags, retained and screen-discarded, blinded to note authorship, identification source, the pipeline's verdict and severity tier. Agreement used Gwet's AC1; proportions carry Wilson intervals. Three propositions were pre-specified: envelope parity, non-differential behaviour across arms, and concordance on consensus cases. Results All ten reviewers completed: 565 adjudications across 434 items, 131 of them double-rated. Inter-clinician agreement on genuineness was fair (raw 59%, 95% CI 50 to 67; AC1 0.24), leaving no human consensus to serve as truth. Judge-clinician agreement was 64% (95% CI 60 to 68), overlapping that interval. Behaviour was near-symmetric on contrast-critical metrics: kept-precision 74% for AI against 81% for clinician notes, and severity signed gap +0.06 against -0.09 tiers. One sub-metric was asymmetric: removed-confirmed 56% against 42%, so the screen over-removes more on clinician notes, a direction conservative to the parent contrast. On 77 consensus items the pipeline concurred on 70% (95% CI 59 to 79). Latent-class triangulation placed the genuine-error rate among flagged candidates at 68% (94% credible interval 48 to 83). Conclusions The judges behave as a consistent, near-non-differential, clinician-equivalent instrument. This licenses a directional AI-versus-clinician contrast under a non-differential misclassification argument, subject to its conditions. It is not a claim of accuracy, which moderate consensus concordance and fair reliability preclude, and the genuine-error rate is best reported as an interval.
Keegan, L.; Shoaf, K.
Show abstract
Infectious disease dynamics is a growing, interdisciplinary field that aims to advance the understanding of how infectious diseases spread and how to control them. Most trainees enter the field through established disciplines and assemble ad hoc training and experience in infectious disease dynamics. As such, expectations for doctoral training remain largely implicit and highly variable across institutions. Other fields have formalized training expectations though defined training competencies, which promote transparency and alignment across institutions without prescribing specific approaches to training or research. In this paper, we set out to define the core competencies that characterize doctoral-level expertise in infectious disease dynamics. We assembled a team of seven people at the University of Utah and drafted a competency set. We then validated the competency set with experts in the field using an e-Delphi process. We did not restrict participation by location, job title, or sector. We set an a priori threshold for consensus to 70% and sent out two rounds of surveys to experts, asking them to rank the competencies by order of importance. Our team initially generated a list of 13 proposed Cross-cutting, 24 Applied Modeling, 17 Data Science, and 16 Theory competencies. After completing two rounds of validation, we validated two tracks comprised of 7 Cross-cutting, 10 Applied Modeling, and 12 Theory competencies. This study represents the first structured effort to define doctoral-level competencies in infectious disease that can help guide curriculum development, comprehensive exam preparation, and trainee evaluation, while also supporting alignment between academic training and workforce needs.
Austria, D.; McCollister, B.; Lindsey, J. E.; Arowolo, M.; Okon, M.
Show abstract
Objective. Formal large language model (LLM) evaluations score isolated prompts, but clinicians and health-informatics researchers meet model failures inside multi-step workflows where erroneous output can alter procedures or contaminate documents. We present TRACE (Tracking Reliability of AI-generated Conversational Evidence), a practitioner-audit framework for evaluating the downstream workflow reliability of conversational AI. Materials and Methods. A method paper with an empirical demonstration: 45 documentation-positive incidents recorded by one clinician-informatician across scholarly, clinical informatics, and clinical-adjacent workflows over seven weeks, coded with a consequence-based severity rubric, an error definition, a taxonomy crosswalk, and a Response-Audit Scorecard. Three reviewer-authors independently coded a 16-incident subsample; three vendor-blinded AI comparators applied the taxonomy to all 45 incidents. Results. Four categories tied as most frequent: verification failure, factual numerical error, tool-behavior misunderstanding, and citation or reference formatting (n=7 each). Four workflow-harm patterns recurred: procedural propagation, documentary contamination, trust-calibration disruption, and user-borne corrective burden, and one incident carried an estimated $2500 impact. Category agreement across three human reviewer-authors was low (Fleiss {kappa}=0.155), whereas three AI comparators agreed substantially (Fleiss {kappa}=0.632), suggesting taxonomy legibility under standardized conditions even where human judgment diverged. Discussion. Category assignment is comparatively legible, whereas severity and claimed-verification remain judgment-dependent. The claimed-verification gap is a measurable failure mode distinct from hallucination, sycophancy, and over-refusal. Conclusion. Practitioner audits with structured response scoring complement benchmarks by documenting workflow harm as an applied evaluation unit for clinical informatics and public-health work; this is a pilot that motivates, not estimates, error rates or cross-model comparisons.
Chen, Y.; McMurry, A.; Gottlieb, D.; Jones, J. R.; Strober, B. J.; Mandl, K. D.
Show abstract
Objective. Privacy regulation constrains sharing line-level electronic health records (EHR) across institutions. One alternative is to aggregate counts into a cube, a table of counts for every combination of categorical variables, with cells below a threshold suppressed. This study asked whether common analyses on the cube reproduce conclusions from line-level data, and whether suppression prevents recovery of the small cells it is meant to hide. Materials and Methods. A Bayesian count-inference pipeline was built that reconstructs suppressed counts and doubles as a reconstruction attack. Applied to 285 pediatric kidney-transplant patients at Boston Children's Hospital, statistical fidelity (Jensen-Shannon divergence, Cramer's V, and R2) and analytical utility (marginal distributions, subgroup graft rejection odds ratios, and logistic-regression classification) were evaluated. Conditional Tabular GAN (CTGAN) synthetic data served as a comparator. Results. Statistical analyses on the cube recapitulated results from line-level data. Across 106 demographic-by-medication subgroups, a bootstrap mean of 3.5 subgroups showed a significant graft-rejection association. The cube's odds-ratio sign changes reversed no significant associations, versus 2.3 for CTGAN. The same reconstruction also defeated suppression: in a 10-variable cube, 76.6% of suppressed cube cells were recovered exactly (14,554 of 18,994), including 85.5% of single-patient cells. Discussion. The cube reproduced common kidney-transplant analyses, but the same reconstruction also recovered suppressed cells; fidelity and privacy risk are thus two faces of one reconstruction rather than independent properties. Conclusions. The cube is a useful surrogate for these kidney-transplant analyses only when paired with a stronger privacy mechanism. This study demonstrated reconstructability of suppressed counts, not re-identification.
Alwakeel, M.; Zaveri, S.; Buck, E.; Rajagopal, S.; Verma, D.; Loriaux, D.; Henao, R.; Tapson, V. F.; Ortel, T. L.; Jones, W. S.; Martin, J. G.; Haines, K. L.; Freeman, N. L.; Wong, A.-K. I.
Show abstract
Background: The 2026 American Heart Association/American College of Cardiology (AHA/ACC) guidelines replaced the 2019 European Society of Cardiology (ESC) four-tier pulmonary embolism (PE) risk scheme with five clinical categories (A-E) and subcategories. These categories were set by expert consensus and have not been validated against outcomes. How patients are reclassified relative to ESC, or how the two systems compare prognostically, is unknown. Methods: We utilized three cohorts of patients with confirmed PE using structured electronic health record data, laboratory biomarkers, and large-language-model abstraction of radiology reports: Duke University Health System (n=12,992, drawn from 95,760 consecutive inpatient CT pulmonary angiography studies, 2014-2025, with no referral or registry enrollment step between imaging and cohort entry), INSPECT (Stanford; n=3,870), and MIMIC-IV (Beth Israel Deaconess; n=361). Patients were assigned AHA/ACC categories B through E, subcategorized where data allowed, and mapped to 2019 ESC risk strata. The primary outcome was 30-day mortality; discrimination was assessed with Harrell C-index. Results: Among 17,223 patients with confirmed PE, pooled 30-day mortality rose monotonically across categories: 1.5% (B), 8.9% (C), 15.5% (D), and 31.9% (E), with the ordering preserved in all three cohorts despite differing baseline mortality. Subcategory-level discrimination was reliable only at the high-acuity extreme (D2-E2); across subcategories C1 through D1, mortality did not order monotonically (9.2%, 10.8%, 8.1%, 10.9%), and adding subcategories to category C did not improve discrimination at Duke (C-index 0.699 vs 0.699). Category C patients lacking both echocardiography and biomarker testing (12.7% of category C) had mortality (10.4%) equal to or exceeding classified peers. Relative to ESC, the frameworks were concordant at the extremes, but 5.7%of ESC intermediate-risk patients were reclassified to category D, with modestly higher but non-significant 30-day mortality than those remaining in category C (10.8% versus 8.9%). Conclusions: Across a three-health-system cohort, the 2026 AHA/ACC framework produced a reproducible mortality gradient at the category level, with added subcategory granularity refining risk chiefly at the highest-acuity tiers. Discrimination across the broad intermediate band was limited, and reclassification from ESC fell almost entirely within this range.
Yehoshua, A.; Lupton, L. L.; Hu, T.; Cappelleri, J. C.; Gavaghan, M. B.; Puzniak, L.; Brathwaite, R.; Di Fusco, M.; Sun, X.
Show abstract
Background To characterize Coronavirus disease 2019 (COVID-19) symptom severity, and recovery from pre-infection through one month, overall and by risk groups. Methods Symptomatic adults aged [≥]18 years with test-confirmed COVID-19 were enrolled from ambulatory care clinics within a national U.S. retail pharmacy network between 10/24/2024 and 08/29/2025 (NCT05160636). Adjusted mixed models for repeated measures estimated least-squares mean changes (LSE) and standard errors (SE) from pre-infection and on Days 1-7, 10, 14, and Week 4 from enrollment in composite symptom scores (sum of severity ratings (0-3) across 14 symptoms), counts of mild-to-severe, moderate-to-severe, and severe symptoms, overall and by age and clinical risk status. Effect sizes (ES) were defined as small (0.2-<0.5), medium ([≥]0.5), and large ([≥]0.8). Results The analysis included 608 adults. On Day 1, symptom severity rose sharply from pre-infection for the composite symptom score (LSE 14.2 [SE 0.3]; ES 2.22), mild-to-severe (7.6 [0.1]; 2.72), moderate-to-severe (5.0 [0.2]; 1.77); and severe (1.8 [0.1]; 0.92) (all p<0.001). By Week 4, composite score (0.7 [0.2]; 0.26), mild-to-severe (0.5 [0.1]; 0.23); moderate-to-severe symptoms (0.1 [0.1]; 0.17) and severe symptoms (0.2 [0.1]; 0.5) remained slightly above baseline (all p[≤]0.025). Elevated severe symptom durations varied: high-risk adults (through Day 3), adults <50 years (through Day 7), and adults [≥]50 years (through Day 7). Conclusions COVID-19 was associated with notable acute symptoms in outpatients, followed by gradual improvement over time, although symptoms still persisted at four weeks. Improvement in severe symptoms varied by individual risk profile, reinforcing the importance risk-based follow-up and ongoing monitoring.
Bergman, H. I.; Liu, V.; Austin, B.; Ali, S.; Fiedler, M.; Sandiford, C.; Blanchard, R.; Casanovas, C. L.; Pedrazzini, G.; Markopouliotis, T.; Vermersch, F.
Show abstract
Background Ambient AI documentation tools, known as scribes, are entering routine clinical practice at scale, but the evidence comparing the notes they produce against clinician-written notes is dominated by single-site, single-language studies that rely on human review to find errors, a method known to miss most documentation errors. Methods We conducted a paired simulation across five countries and languages (Cambridge/English, Barcelona/Spanish, Milan/Italian, Paris/French, Cologne/German; 385 paired consultations, 770 notes). From each actor-performed consultation, an AI scribe (Heidi) and a junior-to-middle-grade clinician independently produced a note. Notes were scored on the PDQI-9 by evaluators blinded to authorship. Documentation errors were identified by two methods of deliberately different sensitivity - clinician adjudication, and a calibrated automated reviewer externally validated against a blinded ten-clinician panel - then graded for clinical risk by a three-model panel. The co-primary outcomes were PDQI-9 total and Critical+High error burden, the latter reported under both detection arms. The analysis plan was registered before any pooling across sites. Results AI notes scored higher than clinician notes on the PDQI-9 (40.6 vs 35.6; difference +5.08, 95% CI 4.6-5.6; Cohen dz=0.55), consistently across all five sites (dz 0.41-0.75), and were less dispersed (5.7% of AI vs 27.8% of clinician notes fell below the study pre-specified low-score threshold (<32)). On the principal safety outcome - the paired probability that a note carried [≥]Critical+High error - clinician notes were affected more often under both detection arms: 61.0% versus 24.4% by the calibrated reviewer (relative risk 2.50, 95% CI 2.09-3.00) and 21.8% versus 6.2% by clinician adjudication (relative risk 3.50, 95% CI 2.32-5.27). The difference was largest for omissions. Unaided clinician review identified roughly 12% of the errors the calibrated reviewer retained, and a smaller fraction in AI notes than in clinician notes. Conclusions In this simulation, AI-generated notes scored higher on documentation quality, varied less, and carried fewer clinically significant errors than notes written on the same consultations by junior-to-middle-grade clinicians. The magnitude of the safety difference depends on the sensitivity of error detection, so we report both detection regimes and bound rather than point-estimate the absolute error rate. Extension to live practice, consultant-authored documentation, and notes as filed after clinician editing remains to be established.
Gorenshtein, A.; Omar, M.; Barash, Y.; Kruskal, J. B.; Ahmed, M.; Brook, O. R.; Klang, E.
Show abstract
Clinical AI agents may be assigned to individual patients, but hospital resources are shared across many patients. We tested what agents do when helping their assigned patient would violate the hospital's rule for a scarce resource. We analyzed 22,916 simulated cases comprising 274,992 logged agent actions across 20 AI models. In each scenario, the agent could claim a scarce resource for its patient even though the hospital rule gave another patient priority. We varied only the agent's assigned role, from responsibility for the whole ward to strong advocacy for one patient. Violations of the hospital rule rose from 32.5% under whole-ward responsibility to 69.4% under strong patient advocacy, a 36.9-point increase (95% CI, 25.7-48.0). Agents correctly identified which patient should receive the resource in 95.7% of tests, yet still took it for their own patient in 65.9% of those episodes. Asking the agent to apply its own allocation judgment immediately before acting reduced violations to 0-2% in a three-model follow-up experiment. Assigned roles can shape how clinical AI agents use shared hospital resources, even when they identify the correct priority patient. Patient-focused agents should not independently control shared resources without an allocation check.
Gorenshtein, A.; Jia, E. L.; Omar, M.; Brook, O. R.; Ahmed, M.; Kruskel, J. B.; Barash, Y.; Klang, E.
Show abstract
Safety alignment should persist while a language model performs a task. We tested whether a single-patient triage task suppressed a warning about a second patient. Each case centered on Patient 1; Patient 2's urgent problem appeared only in passing. Sixteen models saw each case twice: once as a general assistant and once while producing a triage record for Patient 1. As general assistants, models warned the caller in 87% of cases; under the task, they did so in 21%. Every model showed a significant decrease. Yet under the task, the record still mentioned Patient 2 in 76% of cases and recommended urgent care in 67%. Across 15 open-weight models, repeating the emergency-care instruction raised the warning rate only to 29%; moving the message-to-caller field to the top raised it to 36%. Current safety alignment did not reliably persist under task assignment.
Dashti, N.; Schneider, M. M. K.; Eckardt, J. N.; Fiebig, F.; Schweigler, D.; Buttner, S.; Middeke, J. M.; Bornhauser, M.; Rollig, C.; Kather, J. N.; Wiest, I. C.
Show abstract
Background: Adverse event (AE) coding is essential for safety monitoring in oncology clinical trials, particularly in acute myeloid leukemia (AML), where intensive therapies are associated with frequent and heterogeneous toxicities requiring standardized MedDRA (Medical Dictionary for Regulatory Activities) coding. However, manual Low-Level Term (LLT) assignment remains labor-intensive, subjective, and difficult to scale. Although large language models (LLMs) have emerged as promising decision-support tools for automated coding, unguided zero-shot generation remains insufficient for reliable fine-grained MedDRA coding. Objective: To develop and evaluate a retrieval-augmented reasoning pipeline for clinically aligned LLT-level MedDRA coding of free-text adverse events from prospective AML clinical trials. Methods: We implemented a retrieval-augmented reasoning pipeline inspired by the retrieval-augmented generation (RAG) paradigm using LLaMA-3.3-70B-Instruct as the primary backbone and benchmarked the framework across multiple open instruction-tuned LLMs. Dense semantic retrieval first generated a constrained top-100 LLT candidate set for each AE, followed by structured LLM reasoning to select a single best-matching LLT and deterministic mapping to Preferred Term (PT) and System Organ Class (SOC) levels. The pipeline was evaluated retrospectively on AE datasets from three prospective AML clinical trials (MOSAIC, DELTA, and DaunoDouble) with automated LLT/PT/SOC metrics and expert-assessed Clinical Correctness Rate (CCR). Results: Clinical expert review showed high clinical acceptability of the RAG pipeline across datasets (91-97%). Under automated evaluation, the pipeline achieved LLT exact accuracy of 50-58%, PT accuracy of 78-85%, and SOC accuracy of 90-93%. Zero-shot generation and random candidate selection performed substantially worse. Semantic retrieval more often included the coder-assigned LLT among the candidate terms available to the model than retrieval based on lexical similarity. Multi-model benchmarking showed that backbone choice mainly affected LLT exact agreement, whereas PT and SOC performance remained comparatively stable. Conclusions: Retrieval-augmented reasoning supports clinically aligned MedDRA coding of free-text adverse events under realistic candidate constraints in AML clinical trials. Evaluation across three AML clinical trials showed that strict LLT-level string agreement underestimated clinical ap-propriateness, highlighting the importance of combining hierarchical evaluation metrics with clini-cal expert validation for AI-assisted MedDRA coding in hematology trials.
Waterfield, T.; Taylor Miller, P.; McDowell, C.; Agus, A.; Murphy, L.; Sanders, C.; Kearney, A.; Sherrett, F.; Wyche, J.; Hartshorn, S.; Bandi, S.; Blackwood, B.; Williams, N.; Roland, D.; Ferris, K.; Marshall, A.; Clarke, M.; Sutcliffe, A.; Woolfall, K.
Show abstract
Background Obtaining uncontaminated urine samples from children can be difficult. Clean catch urine (CCU) is non-invasive but may be slow and lead to a contaminated sample, whereas transurethral bladder catheterisation (TUBC) and suprapubic aspiration (SPA) are invasive. We assessed the feasibility of randomising children to a definitive trial. Methods FROG was a multicentre, randomised feasibility trial with a mixed-methods perspectives study, health-economic analysis and stakeholder consensus meeting. Children under 16 years requiring urine testing for suspected urinary tract infection (UTI) who could not provide a midstream sample were eligible for the feasibility trial. Parents, children and healthcare professionals were eligible for the perspectives study and consensus meeting. Results Of 703 children screened, 170 were offered the study and 99 were recruited. Overall, 64/170 (37.6%) consented to randomisation, exceeding the feasibility threshold (33%); 32 were allocated to CCU and 32 to TUBC. The allocated method was received by 46/64 (71.9%); delays, unsuccessful collection and distress contributed to non-receipt. Among participants with available cultures, contamination occurred in 2/12 (16.7%) allocated CCU and 0/6 allocated TUBC. No participants consented to randomisation involving SPA. The perspectives study included 14 parent interviews, 89 parent questionnaires and 28 staff across 5 focus groups and 1 interview. CCU and TUBC were considered acceptable, although participants balanced speed and accuracy against pain and distress. SPA availability and acceptability were limited. A total of 19 stakeholders attended the consensus meeting; 94% supported recruiting children aged under 18 months and 100% supported comparing CCU with TUBC, without SPA. Accuracy was the highest-ranked outcome. Conclusions A definitive trial comparing CCU-first with TUBC-first in children aged under 18 months is feasible. Its primary outcomes should reflect diagnostic accuracy and clinical consequences of contamination, with successful collection, collection time, pain and distress assessed as key secondary outcomes.
Patel, K.; Pan, T.; Al-Kindi, S.; Eagar, T. N.; Torre-Amione, G.; Guha, A.; Ranka, R.; Gao, R.; Bhimaraj, A.
Show abstract
BACKGROUND: Increased left ventricular mass (LVM) at a single time point after heart transplantation (HT) predicts future adverse outcomes. However, dynamic changes in LVM could have better biological relevance and reflect adverse graft remodeling (AGR). The prognostic significance of such serial changes has not been studied. METHODS: Using an automated, electronic health record-based institutional data infrastructure, we studied 439 HT recipients with 5,563 LVM measurements. Separate Bayesian joint models estimated the simultaneous associations of current LVM and its instantaneous rate of change with graft dysfunction (GD) and mortality. A joint-model-derived remodeling score combining patient-specific deviations in LVM and slope was dichotomized to define AGR and non-AGR groups. A mixed-effects analysis of all clinical variables was performed to assess associations with LVM both between and within patients. An independent cohort of 35 patients with 79 surveillance-biopsy RNA-sequencing samples was used to examine early stress-responsive pathways associated with the remodeling score. RESULTS: LVM declined by approximately 7 g/year after transplantation, with regression attenuating over time. Sixty patients (13.7%) had GD, and 75 (17.1%) died. Higher LVM was associated with subsequent GD (hazard ratio [HR] per 10 g, 1.14; 95% credible interval [CrI], 1.02-1.28) and mortality (HR, 1.10; 95% CrI, 1.02-1.19). A more positive LVM slope was associated with GD (HR per 1 g/year, 1.21; 95% CrI, 1.06-1.42) and with cardiac allograft vasculopathy (CAV) grade 2 or 3 (HR, 1.39; 95% Crl, 1.02-1.96). LVM regressed more slowly in the AGR group (-5.8 vs -8.4 g/year), with higher GD (21.0% vs 6.4%) and mortality (24.2% vs 10.0%). Time-updated GD was associated with subsequent death (HR, 8.12; 95% Confidence Interval [CI], 4.67-14.14). Transcriptomic analysis showed enrichment of interferon-mediated signaling and vascular endothelial activation with higher remodeling scores, whereas lower scores were associated with mitochondrial and metabolic processes, ribosome biogenesis, and pathways related to tissue repair and stress responses. CONCLUSIONS: AGR is an easily accessible imaging biomarker that reflects the changes in the allograft in response to various stressors and predicts future adverse outcomes. Discovery of molecular mechanisms of AGR could lead to novel therapies to protect the allograft from chronic rejection.
Pybus, A.; Qiu, J.; Morais Lyra, P. C.; Dang, K.; Narvaez-Bandera, I.; Jolaogun, T.; Goecks, J.
Show abstract
Survival analysis is a fundamental technique in biomedical research for modeling time-to-event data. It enables the identification of prognostic factors in disease, compares survival outcomes across treatment groups, and performs targeted treatment selection. A variety of machine learning (ML) approaches to survival analysis have emerged to complement classical statistical methods, especially for high-dimensional datasets with complex, nonlinear interactions between features. However, using survival ML methods requires addressing challenges such as censoring-unaware evaluation, overfitting, selecting performance metrics, and data leakage. To address these and other difficulties in using survival ML models, we developed the mlsurv software package. mlsurv is an open-source Python package built around three major design principles: 1) methodological rigor, including evidence-based model selection, leakage-free pipelines, and multi-metric evaluation, 2) multi-scale evaluation and interpretation, including population and subpopulation evaluation, patient-level explanations, and feature analysis, and 3) automated trust and transparency, including limitation flagging and TRIPOD+AI-aligned reporting. mlsurv bundles ten models spanning linear, ensemble, kernel, and deep learning families within a unified software package. We demonstrate mlsurv on the Chowell immunotherapy cohort (n=1,479). The survival-trained models achieve a test concordance index of 0.73 for overall survival prediction. Further, risk scores strongly correlate with the response-trained LORIS clinical score (|{rho}| up to 0.84), reflecting the overlap between prognostic and predictive signal. mlsurv enables biomedical researchers to conduct rigorous, multi-model survival analysis and benchmarking using minimal code with default best practices rather than implementing custom scripts and methodological safeguards from scratch.